Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records
نویسندگان
چکیده
BACKGROUND Cardiovascular disease (CVD) has become the leading cause of death in China, and most of the cases can be prevented by controlling risk factors. The goal of this study was to build a corpus of CVD risk factor annotations based on Chinese electronic medical records (CEMRs). This corpus is intended to be used to develop a risk factor information extraction system that, in turn, can be applied as a foundation for the further study of the progress of risk factors and CVD. RESULTS We designed a light annotation task to capture CVD risk factors with indicators, temporal attributes and assertions that were explicitly or implicitly displayed in the records. The task included: 1) preparing data; 2) creating guidelines for capturing annotations (these were created with the help of clinicians); 3) proposing an annotation method including building the guidelines draft, training the annotators and updating the guidelines, and corpus construction. Meanwhile, we proposed some creative annotation guidelines: (1) the under-threshold medical examination values were annotated for our purpose of studying the progress of risk factors and CVD; (2) possible and negative risk factors were concerned for the same reason, and we created assertions for annotations; (3) we added four temporal attributes to CVD risk factors in CEMRs for constructing long term variations. Then, a risk factor annotated corpus based on de-identified discharge summaries and progress notes from 600 patients was developed. Built with the help of clinicians, this corpus has an inter-annotator agreement (IAA) F1-measure of 0.968, indicating a high reliability. CONCLUSION To the best of our knowledge, this is the first annotated corpus concerning CVD risk factors in CEMRs and the guidelines for capturing CVD risk factor annotations from CEMRs were proposed. The obtained document-level annotations can be applied in future studies to monitor risk factors and CVD over the long term.
منابع مشابه
A Bootstrapping Approach to Symptom Entity Extraction on Chinese Electronic Medical Records
Symptom entities are widely distributed in Chinese electronic medical records. Previous approaches on symptom entity extraction usually extract continuous strings as symptom entities and require massive human efforts on corpus annotation. We describe the symptom entity as two-tuples of and design a soft pattern matching method to locate them in sentences in the EMR. Our bootst...
متن کاملUsing big data to improve cardiovascular care and outcomes in China: a protocol for the CHinese Electronic health Records Research in Yinzhou (CHERRY) Study
INTRODUCTION Data based on electronic health records (EHRs) are rich with individual-level longitudinal measurement information and are becoming an increasingly common data source for clinical risk prediction worldwide. However, few EHR-based cohort studies are available in China. Harnessing EHRs for research requires a full understanding of data linkages, management, and data quality in large ...
متن کاملThe comparison of cardiovascular risk scores using two methods of substituting missing risk factor data in patient medical records.
BACKGROUND Targeted screening for cardiovascular disease (CVD) can be carried out using existing data from patient medical records. However, electronic medical records in UK general practice contain missing risk factor data for which values must be estimated to produce risk scores. OBJECTIVE To compare two methods of substituting missing risk factor data; multiple imputation and the use of d...
متن کاملTrends in the Prevalence of Diabetes Mellitus in Patients with Myocardial Infarction in the South of Iran: 2008 to 2014
Background: Diabetes mellitus is a strong risk factor for cardiovascular disease, including acute myocardial infarction (AMI). Management of risk factors and the other prevention services in recent years lead to a significant decrease in AMI incidence. However, to examine the success of those strategies to control diabetes, this study aimed to identify the trends in prevalence of diabetes melli...
متن کاملSubtypes of Benign Breast Disease as a Risk Factor of Breast Cancer: A Systematic Review and Meta Analyses
Background: Researchers suggest that benign breast disease (BBD) is a key risk factor for breast cancer. The present study aimed to determinate the risk level of breast cancer in terms of various BBD subgroups.Methods: A meta-analysis was performed to determinate the risk of breast cancer associated with BBD. Observational studies (traditional case-control studies, nested case-control studies, ...
متن کامل